Phase 2: Data & Mathematics Lesson 4 of 5

Python for AI:
Your First Notebook

This is where the theory becomes real. No installation, no setup headaches. Just open a browser, go to Google Colab, and start writing code. By the end of this lesson, you will have loaded a real dataset, asked it questions, and plotted your first chart.

You will learn
Set up and navigate Google Colab in under 2 minutes
Python fundamentals: variables, lists, loops, functions
Load and explore a dataset using Pandas
Create your first data visualisation with Matplotlib

Why Python? Why Colab?

Python is the language of AI. Not because it is the fastest, and not because it was designed for AI. It was not. Python is dominant in AI because of its extraordinary ecosystem of libraries: NumPy, Pandas, Scikit-learn, TensorFlow, PyTorch. Decades of brilliant people have built tools that make incredibly complex operations available in just a few lines of readable code.

Google Colab is a free, cloud-based environment that runs Python notebooks in your browser. No downloads, no installation, no "it works on my machine" problems. You get access to free GPUs, built-in libraries, and the ability to share your work with a link. It is used by researchers at DeepMind and students in their first week of learning AI alike.

1
Open Colab
Go to colab.research.google.com. Sign in with a Google account. Click "New notebook." You are in.
2
Understand the interface
A notebook is a series of cells. A code cell runs Python. A text cell is for notes (uses Markdown). Press Shift + Enter to run a cell.
3
Run your first cell
Type print("Hello, AI world!") in a code cell and press Shift + Enter. You are now a Python programmer.

The four libraries you need

You will work with four core libraries throughout this course. They are all pre-installed in Colab. This is what each one does.

NumPy
Numbers at speed
Fast operations on arrays and matrices. The foundation everything else is built on. import numpy as np
Pandas
Data in tables
Load, clean and analyse datasets as DataFrames, which are essentially programmable spreadsheets. import pandas as pd
Matplotlib
Charts and plots
Visualise data with histograms, scatter plots, line charts. Always the first tool for understanding a new dataset. import matplotlib.pyplot as plt
Scikit-learn
Machine learning models
Train classifiers, regressors and clustering algorithms with a few lines. You will use this heavily in Phase 3. import sklearn

Python fundamentals in 5 minutes

Before we touch data, you need to know a few Python basics. These are genuinely the only things you need to get started. Do not try to memorise them. Just read them once, then use them.

PythonVariables & types
# Variables — no need to declare a type
name = "Alice"
age = 25
is_enrolled = True
score = 94.5

# Lists — ordered collections (can mix types)
ages = [22, 38, 26, 35, 28]
names = ["Alice", "Bob", "Charlie"]

# Access elements (index starts at 0)
print(ages[0])    # 22
print(ages[-1])   # 28 — last element
PythonLoops & conditions
# For loop — iterate over a list
for age in ages:
    print(f"Age: {age}")

# If / else
if age > 30:
    print("Over 30")
elif age == 30:
    print("Exactly 30")
else:
    print("Under 30")

# List comprehension — a compact loop
doubled = [x * 2 for x in ages]
# [44, 76, 52, 70, 56]
PythonFunctions
# Define a function with def
def calculate_mean(numbers):
    return sum(numbers) / len(numbers)

mean_age = calculate_mean(ages)
print(mean_age)   # 29.8

Loading your first real dataset with Pandas

Now for the part that makes everything feel real. The Titanic dataset is one of the most famous datasets in machine learning. It contains information about 891 passengers, including whether each one survived. Here is how to load it and start asking questions.

PythonLoad the Titanic dataset
import pandas as pd

# Load directly from a URL — no download needed
url = "https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv"
df = pd.read_csv(url)

# See the first 5 rows
df.head()
Output
A table appears showing 5 rows × 12 columns: PassengerId, Survived, Pclass, Name, Sex, Age, SibSp, Parch, Ticket, Fare, Cabin, Embarked
PythonAsk basic questions
# Shape: how many rows and columns?
print(df.shape)         # (891, 12)

# Statistical summary
df.describe()

# Average passenger age
print(df['Age'].mean())   # 29.7

# Survival rate (0=died, 1=survived)
print(df['Survived'].mean())  # 0.38 — 38% survived

# Survival rate broken down by gender
print(df.groupby('Sex')['Survived'].mean())
# female    0.742
# male      0.189
What the data is already telling us

With five lines of Pandas, you have already discovered something historically significant: 74% of women survived vs 19% of men. The "women and children first" evacuation policy is visible directly in the data. This is the power of data exploration. Real signals emerge before you build a single model.

Your first plot with Matplotlib

Charts are not just pretty. They reveal patterns in seconds that would take hours to find in raw numbers. The most important habit you can build as an AI practitioner is plotting your data before doing anything else.

PythonPlot passenger age distribution
import matplotlib.pyplot as plt

# Histogram of passenger ages
plt.figure(figsize=(10, 5))
df['Age'].dropna().hist(bins=30, color='#1a6bc8', edgecolor='white')
plt.title('Age Distribution of Titanic Passengers')
plt.xlabel('Age')
plt.ylabel('Number of Passengers')
plt.show()

# Survival count by class (bar chart)
df.groupby('Pclass')['Survived'].mean().plot(
    kind='bar',
    color=['#2a9d8f', '#e76f51', '#264653'],
    title='Survival Rate by Ticket Class'
)
plt.ylabel('Survival Rate')
plt.xticks(rotation=0)
plt.show()

Every time you call plt.show(), a chart appears directly below the cell in Colab. You will see a bell-shaped age distribution and a clear bar chart showing that 1st class passengers survived at a much higher rate than 3rd class, another stark real-world signal sitting right there in the data.

Think of it this way

A Pandas DataFrame is like a smart spreadsheet that you can write instructions to. Instead of clicking through menus, you write short commands. df['Age'].mean() is just "give me the average of the Age column." Once you learn a few dozen of these commands, you can explore almost any dataset in the world.

Hands-on Activity · Google Colab · Pandas
Titanic Explorer
Open Google Colab and complete this exercise. You are not expected to know everything. Experimentation is the point. Use the code from this lesson as a starting template.
01 Load the Titanic dataset using the URL from this lesson. Display the first 10 rows and print the shape of the dataset.
02 Answer these 5 questions with code: (1) What was the average passenger age? (2) What percentage of passengers survived? (3) How many passengers were in each ticket class? (4) What was the youngest and oldest passenger age? (5) What was the average fare paid by survivors vs non-survivors?
03 Create one chart of your choice. A histogram, bar chart, or anything that reveals something interesting about the data. Add a title and axis labels. Share your notebook link in the group.
Open in Colab →
Your Notes
Studying independently? Write your thoughts or answers below. Notes save automatically to your browser.
Progress
Done with this lesson?
Mark it complete to track your progress.